Papers with cross-modal grounding alignments

1 papers
Vision-and-Language or Vision-for-Language? On Cross-Modal Influence in Multimodal Transformers (2021.emnlp-main)

Copied to clipboard

Challenge: Pretrained vision-and-language BERTs aim to learn representations that combine information from both modalities.
Approach: They propose a diagnostic method based on cross-modal input ablation to assess the extent to which pretrained models integrate cross-module information.
Outcome: The proposed method evaluates the model's performance on the other modality based on inputs from one or both modality.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations